Journal of Clinical Epidemiology
○ Elsevier BV
Preprints posted in the last 30 days, ranked by how well they match Journal of Clinical Epidemiology's content profile, based on 31 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.
Li, S.; Zhang, W.; Xing, X.; Shen, Z.; Wang, Y.; Chen, Z.; Neto, O.; Yu, Y.; Wu, C.; Lin, L.
Show abstract
Background Late-stage cancer incidence is being considered as an earlier endpoint in cancer-screening trials, but its trial-level association with cancer-specific mortality may depend on evidence selection and endpoint harmonization. We evaluated the robustness of this association to source-verified additions. Methods We reconstructed the PubMed corpus underlying a 41-comparison review. Gemini 3.1 Pro Preview was used only to prioritize reports for blinded human reassessment. Reviewers determined eligibility, linked reports from the same trial, harmonized endpoints, and verified comparison-level data. We recalculated unweighted Pearson correlations overall and by cancer type after adding earliest-compatible trial comparisons. Results Among 1209 candidate records, 996 PDFs were assessed. Thirty-three reports absent from the source review were prioritized; 26 were eligible, representing 18 trials, and 8 provided compatible comparisons. Adding these comparisons increased the dataset from 41 to 49 and attenuated the overall correlation from 0.73 (95% confidence interval [CI] = 0.55 to 0.85) to 0.59 (95% CI = 0.37 to 0.75). Updated correlations were 0.49 (95% CI = -0.26 to 0.87) for breast, -0.23 (95% CI = -0.71 to 0.40) for colorectal, and 0.83 (95% CI = 0.54 to 0.95) for lung cancer. One sparse-event comparison influenced the colorectal estimate. Conclusions The overall association was sensitive to evidence composition, and cancer-specific stability varied. Late-stage incidence should be evaluated by cancer type and with prespecified sensitivity analyses for evidence selection and endpoint definitions. Model-assisted prioritization cannot replace human eligibility review, trial reconciliation, and source verification.
LIn, H.; Lyu, J.
Show abstract
BackgroundQuality Control Circle (QCC) reports are often reviewed qualitatively, but reviewer workload and inter-rater variability make large-scale assessment difficult. We evaluated whether multiple large language models (LLMs) could score QCC methodological quality reliably on a designed-anchor benchmark. ObjectiveTo estimate inter-model reliability for QCC quality scoring and to assess whether model scores align with designed synthetic anchors and remain descriptively comparable to a small set of public PMC QCC reports. MethodsWe evaluated 30 synthetic QCC reports and 8 public PMC QCC reports across four primary evaluators (GPT, Gemini, Grok, DeepSeek) and one sensitivity evaluator (Claude); Claude was excluded from the primary panel because it shared the model family used during prompt development. Each synthetic case was scored across eight QCC quality dimensions in three runs per evaluator. We summarized each evaluator by median scores, then estimated ICC(A,1) across the primary panel. We also examined score-based calibration against designed anchors, keyword-assisted defect mention, leave-one-out and k=5 sensitivity, and a descriptive synthetic-versus-PMC distributional plausibility check. ResultsInter-model reliability on the primary k=4 panel was excellent: ICC(A,1) = 0.953 (95% CI 0.944 to 0.962) with 237 pooled case-dimension rows. The pre-specified k=5 sensitivity analysis including Claude was 0.954, and leave-one-out estimates within the primary panel ranged from 0.950 to 0.959. Score-based calibration against designed anchors met the prespecified target in 57/58 trap-affected case-dimension rows (98.3%). Keyword-assisted defect mention was present in 51/58 trap instances (87.9%). The synthetic-versus-PMC comparison was descriptively similar across all eight dimensions, and all dimensions met the predefined descriptive margin check. ConclusionsIn this designed-anchor pilot, multi-model LLM scoring of QCC methodological quality showed high inter-model reliability and stable alignment with synthetic anchor scores. These findings support benchmark feasibility, but they do not establish expert validity, clinical validity, or operational deployment readiness.
Tzimas, G.; Vanghelof, J. C.; Mohammed, A.; Raicu, D. S.; Du, L.; Ernst, M. E.; Warner, E. T.; Chan, A. T.; Ryan, J. C.; Espinoza, S. E.; Murray, A.; Sheets, K.; Tchoua, R. B.; Shah, R. C.
Show abstract
Importance: The ASPREE randomized trial found no overall benefit of low-dose aspirin for disability-free survival among older adults. However, individual estimates in pre-specified subgroups indicated potential benefit among racial and ethnic minoritized participants in the United States (US). Objective: To evaluate whether the effect of low-dose aspirin vs placebo on disability-free survival differed across US Black and Hispanic ASPREE participants using individualized treatment-effect estimation. Design, Setting, and Participants: Post hoc clinical trial analysis of ASPREE, a randomized, double-blind, placebo-controlled clinical trial of daily low-dose aspirin vs placebo. This analysis included US ASPREE participants who self-identified as non-Hispanic Black or Hispanic, were aged 65 years or older, and had complete baseline predictor and outcome data. Interventions: Randomization to daily 100-mg aspirin or placebo. Main Outcomes and Measures: The primary outcome was loss of disability-free survival, defined as death, persistent physical disability, or dementia. Individualized treatment effects were estimated post hoc using a Random Survival Forest X-learner. Heterogeneity was evaluated on the relative scale with Cox proportional hazards models and on the absolute scale with 5-year risk differences. Results: Among 2411 US ASPREE participants, 1270 were included in the Black and Hispanic analytic cohort (897 non-Hispanic Black and 373 Hispanic participants; mean age, 71.8 years). Aspirin was associated with lower risk of disability-free survival loss compared with placebo (hazard ratio [HR], 0.65; 95% CI, 0.45-0.93). In model-derived tertiles, aspirin was associated with lower risk in the greatest predicted-benefit group (HR, 0.36; 95% CI, 0.19-0.71; 5-year absolute risk difference [ARD], -11.1 percentage points; 95% CI, -22.0 to -0.1) but not in the lowest predicted-benefit group (HR, 1.26; 95% CI, 0.70-2.27; ARD, +3.9 percentage points; 95% CI, -5.9 to 13.6). Conclusions and Relevance: In these analyses of US Black and Hispanic ASPREE participants, aspirin effects on disability-free survival appear to be heterogeneous, with benefit concentrated in a subset of participants. Because these findings are from post-hoc models, they should be externally validated before being incorporated into clinical decision-making. Trial Registration: ClinicalTrials.gov Identifier: NCT01038583; https://clinicaltrials.gov/study/NCT01038583
McLean, K. W.; LaBonte, J.; Macaulay, K.; Kassam-Adams, S.
Show abstract
This study documents the derivation and validation of a deterministic algorithm for cause-of-death (COD) ascertainment from longitudinal real-world medical claims data, evaluated against an independent state-level death certificate file. Death certificates are the dominant reference standard in mortality research but carry well-documented limitations, including primary-cause error rates estimated at 20-40\% across empirical studies. A matched analytic cohort of 216,382 individuals (Connecticut death records, 2017--2025, age 25 and above) was constructed after exclusion of mechanism-of-injury cases and removal of ill-defined symptom-code entries from both sources. Concordance between algorithmic and certificate-based COD was assessed through three complementary frameworks: age-stratified positive predictive value (PPV) at the ICD-10-CM chapter level under a full-set concordance scenario; mean absolute rank difference (MARD) for chapters identified by both sources; and analyses of breadth, depth, and code-level specificity of COD reporting. Chapter-level PPV was strongest for individuals aged 55 and above, with all estimates representing conservative lower bounds given the known error rate of the certificate reference standard. The algorithm consistently reported broader and more granular contributing cause profiles than the death certificate, with discordances directionally consistent with the well-documented tendency of certificates to under-report contributing conditions. These findings support the conclusion that algorithmic COD ascertainment from longitudinal claims data is a feasible and scalable alternative to certificate-based attribution and, at population scale, a principled methodology for characterising death certificate error rates beyond what small-sample chart review studies can achieve.
MUTHUKA, J. K.; Nyambura, L. W.; Onyango, C. K.; Oluoch, K.; Kioko, M.; Maluki, J.; Nzioki, J. M.; Kim, S.
Show abstract
Background: Autism spectrum disorder (ASD) is a lifelong neurodevelopmental condition for which timely diagnosis is critical to early intervention, family support, and equitable access to care. However, substantial disparities in access to ASD diagnostic services persist across socioeconomic, geographic, clinical, and health-system contexts. This systematic review and meta-analysis synthesized evidence on determinants of access across the ASD diagnostic pathway, from recognition and referral to diagnostic completion and timely diagnosis. Methods: We systematically searched MEDLINE/PubMed, Embase, Scopus, Web of Science, Global Health, and grey-literature sources for studies published between January 2004 and December 2024. Eligible studies examined determinants of ASD diagnostic completion, diagnostic pathways, diagnostic timeliness, or barriers and facilitators to diagnostic access. Two reviewers independently extracted data and assessed methodological quality using the Mixed Methods Appraisal Tool (MMAT). Quantitatively comparable estimates were synthesized using random-effects models with restricted maximum likelihood estimation. Heterogeneity was assessed using Cochran's Q, I2, tau2, and 95% prediction intervals. Pre-specified subgroup analyses, meta-regression, sensitivity analyses, funnel-plot assessments, and Bayesian random-effects analyses were undertaken. Results: The search identified 4,899 records; after removal of 537 records without associated data, 4,362 records underwent title/abstract screening. 3,800 records were excluded, 562 reports were sought for retrieval, and 450 full-text reports were assessed after 112 could not be retrieved. Ultimately, 22 unique studies met the inclusion criteria. Nine unique studies contributed 23 quantitative effect estimates, while the remaining studies contributed to the narrative synthesis. The evidence covered socioeconomic, geographic, family, communication, screening, child developmental, provider, and health-system determinants. The overall random-effects meta-analysis yielded a pooled diagnostic access outcome of 74.1% (95% CI 65.8-81.1%), with substantial heterogeneity (Qe=209.95, p<0.001; I2=88.4%, 95% CI 79.1-94.4%; tau2=0.691) and a wide 95% prediction interval of 32.8-94.4%. Bayesian analysis produced a highly concordant pooled estimate of 73.3% (95% CrI 65.3-80.2%), with I2=87.5% and tau=0.833, and satisfactory MCMC convergence (R-hat=1.000). By outcome domain, pooled successful outcomes were highest for diagnostic pathways (89.3%, 95% CI 70.1-96.7%), followed by timely diagnosis (76.3%, 95% CI 62.9-86.0%), and lowest for diagnostic completion (67.1%, 95% CI 61.8-72.0%) (Qm=5.98, p=0.050). Timely diagnosis demonstrated particularly high heterogeneity (I2=91.2%), whereas diagnostic completion showed moderate heterogeneity (I2=40.6%). Across determinant domains, frequentist pooled estimates were 79.5% for child developmental/neurobehavioral factors, 74.2% for family/socioeconomic/perceptual factors, 68.0% for intervention/care-navigation factors, and 63.6% for provider/clinical recognition factors. Bayesian estimates were 76.7% (BF=53.76), 72.9% (BF=226.32), 64.3% (BF=25.60), and 53.7% (BF=0.684), respectively. Meta-regression indicated that determinant category (Qm=13.48, p=0.004) and effect measure (Qm=7.81, p=0.020) significantly explained between-study variation, whereas age group (p=0.203) and geographic region (p=0.453) did not. Family/socioeconomic factors had significantly larger effect sizes (B=2.703, 95% CI 0.661-4.744; p=0.009), as did child developmental/neurobehavioral factors (B=1.516, 95% CI 0.047-2.985; p=0.043). Potential small-study effects were detected by two of three asymmetry tests, although the Rosenthal fail-safe N was 1,723. Trim-and-fill identified seven potentially missing estimates, with an adjusted pooled effect of 68.4% (95% CI 27.7-109.1%). Importantly, exclusion of two influential outlying estimates produced a pooled outcome of 77.1% (95% CI 71.6-81.9%), indicating that the principal finding was robust. Conclusions: Approximately three-quarters of observed ASD diagnostic outcomes represented successful access, but the substantial heterogeneity indicates that diagnostic access is highly context-dependent. Families were more likely to successfully navigate diagnostic pathways than to complete diagnostic assessment, while timely diagnosis showed the greatest variability across settings. Family and socioeconomic circumstances and child developmental characteristics emerged as particularly important determinants, whereas provider-related effects were more heterogeneous and uncertain. Improving equitable ASD diagnosis requires interventions spanning the entire diagnostic pathway, including developmental surveillance, screening, referral coordination, family navigation, provider capacity, specialist availability, and mechanisms to ensure completion of diagnostic assessment. Greater longitudinal and implementation research is particularly needed in low- and middle-income countries, where diagnostic infrastructure and specialist capacity remain limited.
Bunning, B. J.; Weng, Y.; Wu, D. J.; Hui, G.; Hope, J. E.; Pandurangan, V.; Lopez, I.; Everett, S.; Chen, J. H.; Desai, M.
Show abstract
Doctors increasingly rely on AI in the clinic, yet which report features make AI-generated responses useful and trustworthy remains unclear. In this randomized mixed-methods study, 34 oncology physicians provided 294 ratings of four blinded AI systems across five vignettes, alongside 20 semi-structured interviews analyzed with a prespecified LLM-assisted qualitative pipeline. Despite similar references, an evidence-graded report adapted from OpenEvidence was rated significantly lower in overall utility than standard OpenEvidence (mean difference, -0.96; 95% CI, -1.26 to -0.66; P<.001). Qualitative analysis identified six themes and seven design requirements. Oncologists valued rapid orientation, evidence retrieval, and verification, preferring concise, scannable reports with quantitative outcomes, recognizable bolded guidelines, explicit uncertainty, and verifiable citations. Trust deteriorated with citation mismatch, buried provenance, evidence misclassification, overconfident recommendations, and poor organization. Evidence presented differently can alter perceptions of clinical utility and trust; accuracy alone is insufficient, and report design must also be empirically evaluated.
Ndabashinze, R.; Franzen, D.; Kozuch, E.; Aagerup, J.; Fink, A.; Yerunkar, S. S.; Hunter, K.; Mayo-Wilson, E.; Ying, X.; Kilicoglu, H.; Schorr, S. G.; Seidler, A. L.
Show abstract
Background Clinical trials conducted in Germany are registered across multiple registries, including the German Clinical Trials Register (DRKS), ClinicalTrials.gov, the EU Clinical Trials Register (EUCTR), and, since 2023, the Clinical Trials Information System (CTIS). These registries record health conditions using different classification systems and terminologies, including ICD-10-GM, MeSH, MedDRA, and free text, making cross-registry analyses difficult. We developed and evaluated a pipeline for harmonizing trial condition descriptions to WHO ICD-10 and compared its performance with that of a large language model (LLM) and to health conditions coded by humans. Methods We developed a four-stage, registry-aware mapping pipeline consisting of: (i) condition mention extraction and normalization; (ii) classification of ICD-mappable versus non-mappable mentions; (iii) ontology-based candidate generation using UMLS links between MeSH, MedDRA, ICD-10-GM, and WHO ICD-10; and (iv) SapBERT-based semantic retrieval with hybrid confidence scoring. A second variant additionally applied cross-encoder reranking of the top candidate codes. A stratified sample of 500 condition mentions was manually coded to create an expert reference standard. GPT-4o was evaluated in parallel using the same structured decision framework as the human reviewers. Performance was assessed using accuracy, precision, F1 score, and Cohen's {kappa} at the three-character, block, and chapter levels of ICD-10. Results The pipeline was applied to 23,061 clinical trials and identified 39,512 ICD-mappable condition mentions, of which 72.4% received a high-confidence assignment. Against 390 expert-coded mentions, the baseline pipeline achieved 49.0% accuracy at the three-character ICD-10 level ({kappa} = 0.487), increasing to 58.7% at the chapter level ({kappa} = 0.561). The cross-encoder method produced small but consistent improvements across all evaluation levels. Candidate-recall analysis showed that the correct code was present in the retrieved candidate set in only 73.7% of cases. The LLM substantially outperformed both pipeline variants, achieving 96.7% accuracy and near-perfect agreement with expert coding ({kappa} = 0.966) at the three-character level. The LLM also assigned clinically plausible codes to 82.4% of rejected mentions, 62.8% of Tier-3 exclusions, and 92.3% of review-band mentions. Conclusion Automated harmonization of clinical trial condition data across heterogeneous registries is feasible and supports the use of a common ICD-10 framework for cross-registry analyses. The LLMs achieved high agreement with expert coding, and performed better than the deterministic ontology and embedding pipeline, which achieved moderate agreement. These findings indicate that LLMs can support analyses of the distribution of health conditions investigated in clinical trials in Germany.They are a promising tool for classification of other non-standardised trial characteristics in registries. Keywords: Clinical trial registries; ICD-10; disease harmonization; UMLS; entity linking; SapBERT; large language models; clinical research; natural language processing.
Kremer, P.; Schlicker, N.; Hasnaj, R.; Bamberger, J.; Witte, T.; Haase, I.; Mayr, A.; Schmidt, C.; Osteras, N.; Baraliakos, X.; Kuhn, S.; Krusche, M.; Knitza, J.
Show abstract
Objectives To evaluate whether access to a certified large language model (LLM)-based clinical decision support system improves physician diagnostic performance in rheumatology compared with conventional diagnostic resources alone. Methods In this multicentre, open-label, randomised controlled trial, 82 physicians from seven hospitals in two countries were randomised 1:1 to conventional diagnostic resources plus Prof. Valmed or conventional resources alone. Participants assessed three rheumatology vignettes before and after assistance. The primary outcome was top-1 diagnostic accuracy. Secondary outcomes included top-3 accuracy, diagnostic reasoning, confidence, case-processing time and perceived support quality. Results Top-1 accuracy increased from 22.2% to 33.3% in the intervention group and from 23.3% to 35.0% in the control group, with no between-group difference in improvement (adjusted OR 0.99, 95% CI 0.45 to 2.19; p=0.979). Differences in top-3 accuracy, diagnostic reasoning and confidence were also not significant. Assisted case-processing time was substantially shorter with LLM support (94 vs 206 s; adjusted mean difference -112 s, 95% CI -141 to -83; p<0.001). Information timeliness and perceived diagnostic support quality were rated significantly higher in the intervention group. Exploratory analyses showed persistent overconfidence and substantial AI over-reliance. Conclusions Certified LLM-based diagnostic support did not improve diagnostic accuracy compared with conventional resources, but substantially reduced case-processing time and improved perceived support quality. These findings suggest potential workflow benefits while highlighting overconfidence and over-reliance as important safety considerations.
Bruns, N.; Wessel, A.; Biedermann, R.; Fiedler, K. M.; Goretzki, S. C.; Greve, S.; Hannes, T.; Felderhoff-Mueser, U.; Heimann, K.; Mand, N.; Masjosthusmann, K.; Merker, M.; Soler Wenglein, J.; van den Heuvel, I. A.; Westhoff, J. H.; Tsaka, S.; Lieftuechter, V.; Haertel, C.; Dohna-Schwake, C.; Hojeij, R.
Show abstract
Purpose: Outcome consequences of critically ill children treated outside of pediatric intensive care units (PICU) are unknown. We assessed case fatality of children receiving complex intensive care treatment (CICT) by treating department in Germany and explored reasons for admission to adult intensive care units (AICU). Methods: Retrospective study using the German nationwide hospital discharge dataset 2016 to 2023. Cases aged [≥] 28 days and < 18 years receiving CICT were classified as PICU, AICU, or interdisciplinary by department codes. Odds ratios (OR) for in-hospital case fatality were estimated in generalized linear mixed models with the hospital as random effect, adjusted for age, acute organ dysfunction, and chronic conditions. Excess deaths were estimated and a survey among pediatric and adult intensivists was analyzed qualitatively. Results: Of 143,034 cases, 67.8 % were treated in PICUs, 14.0 % in AICUs, and 18.2 % were interdisciplinary. The crude OR for death in PICUs versus AICUs was 1.14 (95 % CI 1.03 to 1.26), reversing to 0.73 (0.63 to 0.84) after adjustment. For PICU and interdisciplinary cases combined versus AICU, the fully adjusted OR was 0.61 (0.54 to 0.70). Estimated excess deaths across the study period were 100, rising to 191 when interdisciplinary cases counted as pediatric. Capacity constraints, organizational factors, and clinical expertise were the main domains underlying AICU admissions. Conclusions: Children treated outside of PICUs had higher risk-adjusted case fatality, while crude figures pointed in the opposite direction. The findings support treating critically ill children in settings with routine pediatric intensive care experience.
Bergman, H. I.; Liu, V.; Austin, B.; Ali, S.; Fiedler, M.; Sandiford, C.; Blanchard, R.; Casanovas, C. L.; Pedrazzini, G.; Markopouliotis, T.; Vermersch, F.
Show abstract
Background Ambient AI documentation tools, known as scribes, are entering routine clinical practice at scale, but the evidence comparing the notes they produce against clinician-written notes is dominated by single-site, single-language studies that rely on human review to find errors, a method known to miss most documentation errors. Methods We conducted a paired simulation across five countries and languages (Cambridge/English, Barcelona/Spanish, Milan/Italian, Paris/French, Cologne/German; 385 paired consultations, 770 notes). From each actor-performed consultation, an AI scribe (Heidi) and a junior-to-middle-grade clinician independently produced a note. Notes were scored on the PDQI-9 by evaluators blinded to authorship. Documentation errors were identified by two methods of deliberately different sensitivity - clinician adjudication, and a calibrated automated reviewer externally validated against a blinded ten-clinician panel - then graded for clinical risk by a three-model panel. The co-primary outcomes were PDQI-9 total and Critical+High error burden, the latter reported under both detection arms. The analysis plan was registered before any pooling across sites. Results AI notes scored higher than clinician notes on the PDQI-9 (40.6 vs 35.6; difference +5.08, 95% CI 4.6-5.6; Cohen dz=0.55), consistently across all five sites (dz 0.41-0.75), and were less dispersed (5.7% of AI vs 27.8% of clinician notes fell below the study pre-specified low-score threshold (<32)). On the principal safety outcome - the paired probability that a note carried [≥]Critical+High error - clinician notes were affected more often under both detection arms: 61.0% versus 24.4% by the calibrated reviewer (relative risk 2.50, 95% CI 2.09-3.00) and 21.8% versus 6.2% by clinician adjudication (relative risk 3.50, 95% CI 2.32-5.27). The difference was largest for omissions. Unaided clinician review identified roughly 12% of the errors the calibrated reviewer retained, and a smaller fraction in AI notes than in clinician notes. Conclusions In this simulation, AI-generated notes scored higher on documentation quality, varied less, and carried fewer clinically significant errors than notes written on the same consultations by junior-to-middle-grade clinicians. The magnitude of the safety difference depends on the sensitivity of error detection, so we report both detection regimes and bound rather than point-estimate the absolute error rate. Extension to live practice, consultant-authored documentation, and notes as filed after clinician editing remains to be established.
Pryymachenko, Y.; Wilson, R.; Abbott, J. H.
Show abstract
Objectives To analyse the long-term effects of a cruciate ligament (CL) injury on health and socioeconomic outcomes. Methods We used a comprehensive national injury insurance database to identify CL injuries occurring in New Zealand between 2009 and 2022, and employed a doubly robust staggered difference-in-differences research design to identify the effects of these injuries on outcomes up to 10 years after injury. The outcomes of interest were healthcare use (hospitalisations, emergency department visits, medications, knee replacement surgery for osteoarthritis), associated healthcare costs, and labour market outcomes (employment rates, income, and government benefit payments). Results We identified 61 344 CL injuries for inclusion in the analysis. Over 10-year follow-up, a CL injury resulted in increased healthcare use (0.6 more hospitalizations [95%CI 0.4 to 0.7], 1.7 more days spent in hospital [95%CI 1.3 to 2.1], 0.4 more emergency department visits [95%CI 0.3 to 0.6], 2.5 more outpatient visits [95%CI 1.8 to 3.2], and 4.7 more medications dispensed [95%CI -1.8 to 11.2]) and public healthcare costs ($7 537; 95%CI 5 888 to 9 186), reduced income (-$6 060; 95%CI -11 644 to -475), and increased benefit payments ($1 152; 95%CI 542 to 1 761). Conclusion CL injuries have long-term impacts on healthcare use and socioeconomic outcomes. Strategies to reduce the incidence of CL injuries have the potential to realise large health and economic benefits.
Rowan, C. G.
Show abstract
Importance: Active pharmacovigilance via sequential target trial emulation can detect adverse drug event (ADE) signals missed by spontaneous reporting, yet signals identified through high-dimensional screening require rigorous, pre-specified confirmation that addresses residual confounding, outcome heterogeneity, multiplicity, and absolute risk. Objective: To confirm or refute previously detected ADE signals associated with atorvastatin initiation among older adults by applying refined and more homogeneous outcome definitions, expanded family- and component-level exclusions, within-outcome false-discovery-rate control, and probabilistic quantitative bias analysis within a sequential target-trial framework. Design, Setting, and Participants: Confirmatory sequential target trial emulation study using Medicare fee-for-service claims (2017-2019). Eligible participants were statin-naive beneficiaries aged [≥]65 years hospitalized for myocardial infarction or cerebral infarction (primary diagnosis, length of stay [≥]3 days) and discharged home. Up to 14 nested daily trials (Trials 0-13) were constructed beginning on the discharge date, with eligibility, treatment assignment, and follow-up synchronized at each trial origin to eliminate immortal time. Primary analyses stacked all eligible trials; a pre-specified sensitivity analysis restricted inference to Trials 0 and 1, which achieved superior covariate balance (maximum standardized mean difference <0.1). Treatment Strategies: Initiation of atorvastatin (strategy A1) versus initiation of any other new outpatient medication (strategy A2). Strategy A0 (no new medication) was retained only to preserve sequential eligibility. Per-protocol effects were estimated after inverse-probability-of-treatment and inverse-probability-of-censoring weighting, with artificial censoring for treatment deviation (including a 30-day grace period) and death treated as a competing risk in Fine-Gray models. Main Outcomes and Measures: Previously detected signals and more granular, clinically coherent alternatives within the same outcome families (i.e., hemorrhagic events, cardiac valve disorders, musculoskeletal injuries, sensory symptoms, abnormal laboratory findings, and hyperglycemic events), defined by Clinical Classifications Software Refined categories plus independently validated Sentinel or published algorithms. Incident events required absence of relevant baseline codes. Confirmation required (1) within-outcome Benjamini-Hochberg q [≤]0.05 with subdistribution hazard ratio (sHR) >1.0 and (2) both the median and 2.5th percentile of the bias-adjusted sHR remaining >1.0 across 5,000 Monte Carlo draws of probabilistic quantitative bias analysis (confounder-outcome risk ratio 1.25-3.00; prevalence difference 0.05-0.25). Absolute risks, risk differences, and numbers needed to harm (NNH) were reported. Stratified analyses examined time windows (1-30, 31-91, 92-182 days), age, sex, and race. Results: Of 70,130 eligible patients, 39,948 initiated atorvastatin and 19,182 initiated another new medication. After weighting, baseline covariates were closely balanced. Acute hemorrhagic cerebrovascular disease was confirmed overall (sHR 1.43, 95% CI 1.00-2.04; risk difference 0.5%; NNH 205) and more strongly in the first 30 days (sHR 2.20, 1.35-3.58); the association persisted in Trials 0 and 1 (sHR 1.50, 1.02-2.20). Related early intracranial hemorrhage signals were likewise confirmed. Nonrheumatic and unspecified valve disorders were confirmed in days 92-182 (sHR 1.48-1.58), as was cardiac valve intervention overall (sHR 1.74-1.83). Sprains, strains, and related composites were confirmed among men (sHR 1.66-1.94). General sensation/perception symptoms and dizziness were confirmed among non-White patients (sHR 1.40-1.43) but only in the unrestricted trial set. Acute hepatic failure was confirmed overall (sHR 1.61-1.72), and biliary tract disease among women (sHR 1.45-1.49). For every confirmed association the proportion of bias-adjusted draws remaining above the null was 1.00. Multiple prior signals, including prediabetes and acute posthemorrhagic anemia, failed the dual confirmation criteria. Conclusions: Sequential target-trial emulations with refined outcome definitions, within-outcome multiplicity control, restriction to optimally balanced early trials, and probabilistic quantitative bias analysis confirmed several ADE signals associated with atorvastatin initiation in older adults--most notably early hemorrhagic cerebrovascular events, cardiac valve disorders and interventions, musculoskeletal injuries in men, and selected hepatobiliary events--while attenuating others. Absolute excess risks were modest yet clinically relevant in a high-risk post-infarction population. These findings support a two-stage active pharmacovigilance paradigm (signal detection followed by rigorous confirmation) and justify heightened clinical vigilance for the confirmed events, while underscoring the need for external validation in independent populations and data sources.
Dobin, D.; Witmer, A. M.; Sweeney, F.; Ryan, T.; Cimino, A.; Haroz, E. E.; Nestadt, P. S.; Wilcox, H. C.
Show abstract
Importance. Systematic reviews and meta-analyses inform suicide-prevention policy and practice, but broad database searches are difficult to screen manually. This limits capture of upstream interventions, such as economic policies, with indirect effects on suicide. Reliable automated screening could make broader and more comprehensive evidence syntheses feasible. Objective. To develop and validate ScreenAgent, a large language model (LLM) agent for title and abstract screening, and a review-specific method for prospectively estimating screening performance. Design, Setting, and Participants. ScreenAgent was validated internally on a prospective meta-analysis, and externally on two published systematic reviews. The correct include and exclude decisions followed standard systematic-review screening methodology. Exposures. ScreenAgent, an LLM agent returning structured include-or-exclude decisions. Records it marked for inclusion were re-checked by a second, cascade pass using a higher-effort LLM. For the external reviews, the agent's prompt was tuned automatically on a small set of labeled examples. Main Outcomes and Measures. We calculated sensitivity, specificity, workload reduction (the percentage of records removed from human review), and agent-versus-human reliability via Cohen kappa. Sensitivity was estimated by direct comparison (internal) and 5-fold cross-validation (external). Results. In the internal validation, ScreenAgent identified 43 of 44 eligible studies (sensitivity 97.7%; 95% CI, 88.2%-99.6%) with a generic prompt applied without any review-specific optimization, specificity 98.0%, and a measured full-corpus workload reduction of 99.4%. The cost was $855.91 for the full 201,064-record corpus (0.43 US cents per record). Agent-versus-human-consensus agreement exceeded human-versus-human agreement (Cohen kappa 0.75 vs 0.64; percent agreement 97.3% vs 95.4%). For two external validation studies, automatic tuning resulted in a cross-validated sensitivity of 95.9% (95% CI, 90.0%-98.4%) and 97.4% (90.9%-99.3%), with workload reductions of 97.4% and 98.4%. Conclusions and Relevance. Suicide prevention efforts often require rapid consolidation of evidence because of the inherent challenges of single studies trying to prevent rare outcomes. On both internal and external validation sets, ScreenAgent identified nearly all eligible studies with human-level reliability for a fraction of a US cent per record while keeping human reviewers as the final arbiters. By making broad searches feasible and screening performance measurable beforehand, this approach can serve as a transparent methodology to strengthen the speed at which we can inform and advance suicide prevention efforts.
Le Guellec, B.; Bentegeac, R.; Tran, V.-T.; El Homsi, M.; Amouyel, P.; Kuchcinski, G.; Hamroun, A.
Show abstract
Background: Large language models have been proposed to improve patient comprehension of radiology reports. However, whether they improve objective understanding remains unproven. Purpose: To evaluate the effect of appending an LLM-generated lay summary to brain MRI reports on objective and subjective patient comprehension in a randomized controlled trial. Materials and Methods: In this randomized controlled trial, 2,727 adult participants from the ComPaRe e-cohort were randomly assigned to interpret six standardized brain MRI reports for headache, presented either in their native format (control; n = 1,401) or appended with a lay summary generated by an open-weights LLM (Mistral Small 3.2) (intervention; n = 1,326). The primary outcome was objective comprehension, defined as the rate of correct classification of whether the report provided a probable explanation for the headache, with ground truth established by four-radiologist consensus. Secondary outcomes included satisfaction, subjective comprehension, anxiety, and willingness to contact a healthcare professional. Generalized estimating equations accounted for repeated within-participant observations. Results: A total of 2,727 participants (mean age, 52 years +/- 15; 75.2% women) were evaluated. Objective comprehension did not differ between arms (58.3% vs 59.4%; odds ratio (OR) 0.97; 95% CI: 0.90-1.06; P = .54). The intervention significantly improved overall satisfaction (64.9% vs 36.7%; OR 3.26; 95% CI: 2.93-3.64; P < .001) and subjective comprehension (50.3% vs 24.0%; OR 3.17; 95% CI: 2.82-3.56; P < .001). High anxiety was modestly reduced (25.1% vs 26.6%; OR 0.92; P = .037). The effect on objective comprehension varied by report type (P for interaction < .001): summaries improved comprehension of symptom-explaining reports (42.4% vs 37.4%; P < .001) but reduced it for normal reports (72.5% vs 76.6%; P = .001). Conclusion: LLM-generated lay summaries appended to brain MRI reports improved patient satisfaction and subjective comprehension but did not improve objective comprehension, indicating a gap between perceived and actual understanding that should be addressed before clinical integration.
Moe-Byrne, T.; Knapp, P.; Golder, S.
Show abstract
Background People with lower levels of literacy or health literacy may struggle to understand conventional health information. Video animations show promise as information tools, yet it is unclear whether video animations help reduce these inequalities in understanding. This study examined whether the effectiveness of video animations in health settings differs according to level of literacy or health literacy. Methods We drew on trials from a recent systematic review of video animations about healthcare or public health topics for patients or the public. We extracted available data on literacy, health literacy, or proxy indicators. One reviewer extracted data and a second checked all entries. Where possible, we conducted subgroup analyses of low and high literacy levels or interaction meta-analyses comparing low versus high literacy groups; otherwise, results were summarised narratively. Results From 88 eligible trials, we extracted health literacy data for 12. Across nine trials reporting knowledge, animations mostly improved knowledge compared with controls in both lower and higher health literacy groups. Effects on attitudes and behaviours were mixed and often small, with few studies reporting results by health literacy level. Across the subgroup analyses available, there was no consistent evidence of a pooled interaction effect of animations according to low and high literacy groups, but both statistical heterogeneity and small subgroup sizes limited precision of estimates. Across 88 trials, 54 (61%) reported education level, 22 (25%) did not, and 12 (14%) involved children or adolescents likely to have similar education levels. Conclusions Overall, the available data suggest that video animations can improve knowledge outcomes in both lower and higher health literacy groups, but their impact on attitudes and behaviour is less clear. Because literacy was rarely reported or analysed in the trials, it remains uncertain whether animations help to reduce literacy-related inequalities in access to, and use of health information.
Krump, P. A.; Blasingame, M. N.; Koonce, T. Y.; Williams, A. M.; Su, J.; Giuse, N. B.
Show abstract
Background: Large language models (LLMs) that use retrieval-augmented generation (RAG) are increasingly used to answer clinical questions, although the evaluation of these systems remains limited. Building on previous studies conducted by our team, this case report aimed to improve upon this knowledge gap by applying a reusable methodology to compare the performance of eight LLMs that utilize RAG techniques for evidence synthesis. Case Presentation: Eight commercially available RAG LLM tools (OpenEvidence, Undermind, Consensus, SciSpace, Elicit, MediSearch, EvidenceHunt, and Scite) were evaluated using twelve ChatGPT-generated clinical questions on the topics of treatment, etiology, and prognosis. To enable comparison, we prompted ChatGPT to identify all key unique medical concepts from the full set of LLM responses to each question. Concepts were categorized as critical ("must-have") or non-critical ("nice-to-have") for answering the clinical question. Experienced information scientists were consulted at each step for their expertise. Descriptive statistics and Kruskal-Wallis tests were used to compare performance across tools and question categories. No significant differences were found among the eight RAG LLMs in their coverage of "must-have" (p=0.95) or "nice-to-have" (p=0.16) key unique medical concepts, and no single tool consistently captured all identified concepts. Conclusions: These findings suggest that RAG LLMs may be supplementary tools for evidence retrieval and synthesis but cannot, at this time, fully replace comprehensive expert review of the medical literature. The evaluation framework presented here may be a useful model for future comparative assessments of rapidly evolving AI evidence synthesis tools.
Han, F.; Wang, J.; Shi, S.; Jin, M.; Ren, C.
Show abstract
IMPORTANCE: A recent meta-analysis showed that chemoimmunotherapy was associated with improved overall survival (OS) compared with immune checkpoint inhibitor (ICI) monotherapy for programmed death-ligand 1 (PD-L1) tumor proportion score (TPS) [≥] 50% advanced non-small-cell lung cancer (NSCLC). However, whether this benefit reflects chemotherapy effect or ICI heterogeneity remains unclear. OBJECTIVE: To reassess the survival benefit of adding chemotherapy to ICI monotherapy using agent-stratified comparisons anchored to chemotherapy. DATA SOURCES: The 24 phase 3 randomized clinical trials included in the original meta-analysis (search date, August 3, 2025). DATA EXTRACTION AND SYNTHESIS: Hazard ratios (HRs) for OS and progression-free survival (PFS) were extracted from each trial in the original meta-analysis. Two analytic frameworks were used: within-agent comparisons (same ICI in both chemoimmunotherapy and monotherapy) and across-agent comparisons (ICI in one treatment strategy only). For within-agent comparisons, a two-stage random-effects meta-analysis was conducted. In stage 1, ICI-specific HRs for chemoimmunotherapy and ICI monotherapy versus chemotherapy were pooled and their ratio was calculated (RHR = HRchemoimmuno/HRmono; RHR < 1 favors chemoimmunotherapy). The RHRs were pooled in stage 2. For across-agent comparisons, RHR was derived from pooled HRs by treatment strategy. MAIN OUTCOMES AND MEASURES: Endpoints were OS and PFS. RESULTS: In within-agent comparisons (4 ICIs; 13 trials; N = 3252), pooled RHR was 0.94 (95% CI, 0.78-1.13; P = .48; I2 = 0.0%) for OS and 0.85 (95% CI, 0.68-1.06; P = .14; I2 = 0.0%) for PFS. In across-agent comparisons (7 ICIs; 11 trials; N = 2231), RHR favored chemoimmunotherapy for OS (0.68; 95% CI, 0.50-0.92; P = .01) and PFS (0.46; 95% CI, 0.37-0.58; P < .001). In a sensitivity analysis restricted to trials of NCCN-recommended regimens, pooled RHR was 1.02 (95% CI, 0.81-1.28; P = .87) for OS. CONCLUSIONS AND RELEVANCE: In the within-agent comparisons, adding chemotherapy to ICI monotherapy did not improve OS or PFS in patients with PD-L1 TPS [≥] 50% advanced NSCLC. The benefit in the original meta-analysis appears driven by across-ICI heterogeneity. These findings are consistent with ICI monotherapy as a standard first-line option and underscore the need for agent-level stratification in across-trial comparisons.
Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.
Show abstract
We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.
Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.
Show abstract
Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.
Jaber, A.; Hughes, L.; Cameron, A. C.; Quinn, T. J.
Show abstract
Background: Systematic reviews of clinical prediction models increasingly include studies using artificial intelligence (AI) and machine learning (ML) methods alongside traditional multivariable regression approaches. A previously published Excel tool enabled standardised data extraction using the CHARMS checklist and risk of bias assessment using PROBAST. The recent publication of the PROBAST+AI framework, which distinguishes the assessment of model development quality from the assessment of model evaluation risk of bias and assesses applicability in both parts, necessitates an updated digital instrument applicable across prediction modelling methods. Methods: We updated an open-access Excel tool to incorporate the full PROBAST+AI framework. The updated template incorporates structural separation between assessment of model development quality and model evaluation risk of bias, with applicability assessed in both parts. It also incorporates updated signalling questions, including those addressing methodological issues particularly relevant to AI/ML, and automates the generation of summary tables and graphical displays. Results: The updated tool (CHARMS & PROBAST+AI Template) contains 11 worksheets and supports data extraction and appraisal for up to 30 prediction models. Dedicated, linked worksheets enable separate assessment of model development and model evaluation, with Domain 4 distinguishing among Apparent, Internal, and External evaluation settings. Key updates include dedicated assessments for predictor pre-processing, class imbalance handling and recalibration, data leakage prevention, and replication of the full model development pipeline within resampling procedures. Automated sheets dynamically format tables and summary charts covering PROBAST+AI parts. Conclusions: The CHARMS & PROBAST+AI Excel template provides a standardised, user-friendly, and rigorous digital framework for systematic reviewers appraising traditional statistical and AI-driven clinical prediction models.